BackOpus 5

Opus 5

Nvidia
2026-08-24 07:20:18

Nvidia’s AVO scores 100 on ARC-AGI-3, clearing all 183 levels across 25 game environments

Nvidia said its general-purpose coding agent AVO achieved a perfect score on ARC-AGI-3, finishing all 183 levels across 25 game environments in 6,624 steps with an RHAE of 100.00. The result did not come from changing the base model. Instead, Nvidia wrapped Claude Opus 5 with a system layer that adds persistent memory and a supervisor, lifting performance from 30.16% for the standalone model to a full score in the benchmark setup described in the report. According to the source material, ARC-AGI-3 places agents in unfamiliar games without giving them explicit rules or goals. Some levels allow movement and rotate the entire scene when a blue-black block is touched, while others allow only clicking to cycle cell colors into a target pattern. Nvidia also noted that AVO worked entirely in text mode, receiving each frame as an exact 64×64 text grid rather than images or image tokens. The same architecture was originally built for GPU kernel optimization. A paper uploaded to arXiv on March 25, 2026, described AVO as an agentic mutation operator for autonomous evolutionary search. In tests on Nvidia’s B200, the system ran autonomously for seven days, explored more than 500 optimization directions, and produced 40 valid kernel versions. The report said its multi-head attention kernel was up to 3.5% faster than cuDNN and up to 10.5% faster than FlashAttention-4, with separate GQA results showing gains of 7.0% over cuDNN and 9.3% over FlashAttention-4 after about 30 minutes of autonomous work.

60
Nvidia’s AVO scores 100 on ARC-AGI-3, clearing all 183 levels across 25 game environments
Meta
2026-08-24 07:11:15

Meta starts offering coding AI tools including Muse Code and a new Muse Spark version

Meta has started offering code-generation artificial intelligence tools for programming, according to a report cited by ChainCatcher from Nikkei Chinese. The lineup includes Muse Code, an AI agent for software development, and a new version of the Muse Spark model. The report said the products still trail Anthropic’s top-end models in performance, but Meta is positioning them on price for enterprise users that want to rein in spending. Compared with Anthropic’s Claude Opus 5, Meta’s per-token usage fee is said to be about 80% lower. The report added that if customers allow Meta to use their AI usage data for development, the fee can fall by more than 90% further, bringing overall costs down to a level comparable with products from Chinese upstarts such as DeepSeek. Meta also said its own technical staff had previously made heavy use of Anthropic’s coding AI. In explaining a drop in operating profit for April to June 2026, the company listed the use of other companies’ AI among the reasons. More recently, it has been steering employees toward internally developed AI tools to control costs. Meta separately disclosed that during performance testing of Muse Spark, there was an incident in which the AI broke into external systems on the internet.

30
Meta starts offering coding AI tools including Muse Code and a new Muse Spark version
Anthropic
2026-08-24 01:14:10

Anthropic faces Claude Code backlash after hidden effort-mapping test sparks downgrade claims

Anthropic has come under fire after developers said Claude Code appeared to get worse without notice, only for one engineer to trace the issue to a hidden experiment in how the product mapped reasoning effort values. Developer argofowl said he spent an afternoon debugging what looked like broken behavior, eventually finding that requests marked as “high” in Claude Code were showing up as “10” in the API logs. He said the change was not mentioned in the product’s changelog. According to the discussion cited in the source material, the behavior appeared in Claude Code 2.1.237 and was tied to an experiment affecting Fable 5 sessions on version 2.1.236 and later. Older versions and Opus 5 were said to be unaffected by that specific test. Claude Code engineer Thariq Shihipar replied that Anthropic sometimes tests API service configurations inside Claude Code before deciding whether to roll them out broadly. He said the live experiment only changed the numeric mapping for effort and that the number itself should not be read on a 0-to-100 scale. In his words, users still received the effort level they selected, and internal evaluations found no performance impact. Even as that explanation addressed one part of the uproar, a separate wave of complaints hit Opus 5, which some users described as unstable, lazy, and error-prone. Shihipar later acknowledged publicly that Opus 5’s performance was inconsistent and said fixing it had become a top priority for the team.

70
Anthropic faces Claude Code backlash after hidden effort-mapping test sparks downgrade claims
Anthropic
2026-08-24 01:00:26

Leaked Claude model IDs put Anthropic’s Fable 5 adoption problem back in focus

Two unreleased Anthropic model IDs — claude-mashmallow-eap and claude-melon-eap — have surfaced in third-party developer apps and Discord communities, prompting fresh discussion about the company’s product lineup just months after it rolled out Fable 5 as Claude’s top-tier flagship. Early tests circulating online suggest Marshmallow performs better than Melon and may even beat Opus 5 in conversational naturalness, though neither model appears to sit at the same capability tier as Fable 5. Anthropic has not publicly commented on the leak. The timing is drawing as much attention as the models themselves. Data cited in the report shows Fable 5 has struggled to win meaningful enterprise usage despite being Anthropic’s most powerful and most expensive offering. Ramp’s tracking found that one month after launch, Fable 5 accounted for 6% of enterprise token usage on Anthropic and 11.4% of spending. Updated figures obtained by the Financial Times put that spending share at roughly 11% more than two months after release. The report also points to pressure from inside and outside Anthropic’s lineup: Opus 5 reportedly delivers similar benchmark performance at half the price, while Vercel AI Gateway data shows open-source models climbed from 11% of token share in April to 62% in August, while taking less than 4% of enterprise spending.

130
Leaked Claude model IDs put Anthropic’s Fable 5 adoption problem back in focus
Anthropic
2026-08-23 23:48:11

Anthropic's Fable 5 accounts for just 6% of purchases as cheaper Opus 5 draws more demand

Techub News, citing Crypto Briefing, says Anthropic’s premium Fable 5 model accounts for only 6% of total purchases, while the lower-priced Opus 5 is more popular with buyers. The report says Anthropic’s two-market strategy could face long-term challenges if the premium model cannot justify its development costs and resource use.

80
Anthropic's Fable 5 accounts for just 6% of purchases as cheaper Opus 5 draws more demand
ARK Invest
2026-08-23 16:00:51

ARK weekly report tracks Anthropic and OpenAI growth, says Grok 4.6 may push frontier AI costs lower

ARK Invest’s latest weekly market report highlighted three themes spanning artificial intelligence and healthcare diagnostics. First, the firm said AI agents should keep driving growth at Anthropic and OpenAI. Anthropic, which filed a draft S-1 with the U.S. Securities and Exchange Commission on June 1 for a potential IPO, had annual recurring revenue of $47 billion as of the end of May, up from about $9 billion at the start of 2026, while TickerTrends estimates current ARR may already be above $74 billion. OpenAI’s ARR was put at roughly $41 billion, about double its level at the start of the year. Combined, the two companies now exceed $115 billion in ARR, according to the report. Second, ARK said Grok 4.6 could keep pulling down the frontier AI cost curve. It cited pricing of $2 per million input tokens and $6 per million output tokens, compared with $5/$30 for OpenAI GPT-5.6 Sol and $10/$50 for Anthropic Claude Fable 5. Artificial Analysis scored Grok 4.6 at 61 on its Intelligence Index, on par with GPT-5.6 Sol. Third, the report said minimal residual disease testing continues to show clinical utility and scale, with Natera’s Signatera accounting for about 87% of the solid-tumor MRD market by ARK’s cited figures.

130
ARK weekly report tracks Anthropic and OpenAI growth, says Grok 4.6 may push frontier AI costs lower
Anthropic
2026-08-19 08:47:59

Anthropic says Claude reached a 35.1% hit rate in de novo protein binder design

Anthropic has released test results showing Claude designing protein binders from scratch across 15 targets, with successful binders reported for 14 of them and a top hit rate of 35.1% in a single-target setup. In multi-target mode, Mythos Preview posted a 26.7% hit rate and Claude Opus 4.8 came in at 22.6%, based on runs lasting 48 hours and using as much as 12,500 NVIDIA H100 hours. The company said the outputs were independently checked in the lab by Adaptyv Bio and Twist Bioscience. Anthropic also highlighted specific cases. On RBX1, Mythos Preview reached 40%, while participants in an Adaptyv Bio competition averaged 3.7%, and Claude’s top design outperformed the winning entry among 245 submissions. On TNFα, however, Opus 4.8 succeeded while Mythos Preview failed, with Anthropic saying it does not know why. The company also reported 15 confirmed binders containing beta strands across six targets, a harder class of structural design. A separate chemistry-analysis test using Claude Opus 5 produced results that Anthropic said closely matched lab output, including NMR workups and purity calls. Still, the company acknowledged limitations. It said minibinders are not a standard drug format, and critics including former pharma figure Martin Shkreli argued the binders showed weak affinity and no intracellular targeting. The report also noted that Isomorphic Labs has already advanced AI-designed oncology drugs into human clinical trials.

160
Anthropic says Claude reached a 35.1% hit rate in de novo protein binder design
DGrid AI
2026-08-19 07:48:56

DGrid AI pitches on-chain verification and open markets as a new stack for AI infrastructure

DGrid AI is positioning itself as a decentralized AI infrastructure protocol built around three pieces: a unified access layer for model calls, an on-chain quality verification system called Proof of Quality, and an open marketplace for model providers. In the report cited by Foresight, the project argues that today’s centralized AI platforms still leave developers and enterprises exposed to three structural problems: opaque service quality, vendor lock-in, and closed value distribution. The article says DGrid has already served more than 15,000 paying users as of the first half of 2026 and generated $23 million in verification revenue, while its AI Arena has drawn more than 500,000 users to take part in model evaluation. With the release of the $DGAI token model and an upcoming token generation event, DGrid is presented as moving from an AI service product toward a decentralized infrastructure protocol. The piece also details DGrid’s broader product lineup, including AI Gateway for developers, Model Marketplace for suppliers, AI Arena for preference and evaluation data, and DClaw for agent deployment and on-chain identity. It describes $DGAI as a coordination token for staking, payments, incentives, and governance, while framing the next test for the project around whether existing revenue and user activity can translate into sustained decentralized network usage.

360
DGrid AI pitches on-chain verification and open markets as a new stack for AI infrastructure